Papers with template-based (synthetic) variants
DeVisE: Towards the Behavioral Testing of Medical Large Language Models (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing evaluations of large language models do not reveal whether their outputs reflect genuine medical reasoning or superficial correlations. |
| Approach: | They propose a framework that probes fine-grained clinical understanding through controlled counterfactuals. |
| Outcome: | The proposed framework is based on demographic and vital signs data from the ICU discharge notes of patients in the intensive care unit (MIMIC-IV). |